Back

Virus Evolution

Oxford University Press (OUP)

All preprints, ranked by how well they match Virus Evolution's content profile, based on 155 papers previously published here. The average preprint has a 0.09% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Delving below the species level to characterize the ecological diversity within the global virome: An exploration of West Nile Virus

Kong, T.; Mei, K.; Wang, A.; Krizanc, D.; Cohan, F. M.

2019-12-13 microbiology 10.1101/2019.12.12.874214 medRxiv
Top 0.1%
77.6%
Show abstract

Efforts to describe the diversity of viruses have largely focused on classifying viruses at the species level. However, substantial ecological diversity, both in virulence level and host range, is known within virus species. Here we demonstrate a proof of concept for easily discovering ecological diversity within a virus species taxon. We have focused on the West Nile Virus to take advantage of its broad host range in nature. We produced a genome-based phylogeny of world diversity of WNV and then used Ecotype Simulation 2 to hypothesize demarcation of genomes into 69 putative ecotypes (ecologically distinct populations), based only on clustering of genome sequences. Then we looked for evidence of ecological divergence among ecotypes based on differences in host bird associations within the Connecticut-New York region. Our results indicated significant heterogeneity among ecotypes for their associations with different bird hosts. Ecological diversity within other zoonotic viruses could be easily discovered using this approach. Opportunities for extending this line of research to human associations of virus ecotypes are limited by missing geographic metadata on human samples.

2
Genetic Information Processing Complexity as a Determinant of Virus Diversity

Pietrokovski, S.; Shaul, Y.

2026-06-03 evolutionary biology 10.64898/2026.06.01.729294 medRxiv
Top 0.1%
76.7%
Show abstract

Viruses exhibit diverse genome architectures and replication strategies that shape their evolutionary trajectories and taxonomic diversification. Here, we test whether the complexity of viral genetic information processing predicts large-scale patterns of viral diversity. We define a propagation index that quantifies a minimal number of steps required for viral genome expression and replication across the Baltimore classes. Using ICTV taxonomy data (1971-2024), we identify a strong consistent linear relationship between the propagation index and viral diversification at both the family and genus levels. This statistically significant association is also observed for DNA and RNA viruses independently. Notably, the correlation persists across decades of ICTV releases despite substantial expansion and restructuring of viral taxonomy. Viruses with simpler propagation strategies consistently exhibit greater diversification, suggesting that genome processing complexity constrains macroevolutionary potential. These findings establish a quantitative link between propagation architecture and viral diversification and provide a predictive framework for understanding large-scale patterns of virus evolution.

3
The evolutionary history of hepaciviruses

Li, Y.; Ghafari, M.; Holbrook, A. J.; Boonen, I.; Amor, N. M. S.; Catalano, S.; Webster, J. P.; Li, Y.; Li, H.; Vergote, V.; Maes, P.; Chong, Y. L.; Laudisoit, A.; Baelo, P.; Ngoy, S.; Mbalitini, S. G.; Gembu, G.-C.; Akawa, P. M.; Gouy de Bellocq, J.; Leirs, H.; Verheyen, E.; Pybus, O. G.; Katzourakis, A.; Alagaili, A.; Gryseels, S.; Li, Y.; Suchard, M. A.; Bletsa, M.; Lemey, P.

2023-06-30 microbiology 10.1101/2023.06.30.547218 medRxiv
Top 0.1%
76.5%
Show abstract

In the search for natural reservoirs of hepatitis C virus (HCV), a broad diversity of non-human viruses within the Hepacivirus genus has been uncovered. However, the evolutionary dynamics that shaped the diversity and timescale of hepaciviruses evolution remain elusive. To gain further insights into the origins and evolution of this genus, we screened a large dataset of wild mammal samples (n = 1,672) from Africa and Asia, and generated 34 full-length hepacivirus genomes. Phylogenetic analysis of these data together with publicly available genomes emphasizes the importance of rodents as hepacivirus hosts and we identify 13 rodent species and 3 rodent genera (in Cricetidae and Muridae families) as novel hosts of hepaciviruses. Through co-phylogenetic analyses, we demonstrate that hepacivirus diversity has been affected by cross-species transmission events against the backdrop of detectable signal of virus-host co-divergence in the deep evolutionary history. Using a Bayesian phylogenetic multidimensional scaling approach, we explore the extent to which host relatedness and geographic distances have structured present-day hepacivirus diversity. Our results provide evidence for a substantial structuring of mammalian hepacivirus diversity by host as well as geography, with a somewhat more irregular diffusion process in geographic space. Finally, using a mechanistic model that accounts for substitution saturation, we provide the first formal estimates of the timescale of hepacivirus evolution and estimate the origin of the genus to be about 22 million years ago. Our results offer a comprehensive overview of the micro- and macroevolutionary processes that have shaped hepacivirus diversity and enhance our understanding of the long-term evolution of the Hepacivirus genus. SignificanceSince the discovery of Hepatitis C virus, the search for animal virus homologues has gained significant traction, opening up new opportunities to study their origins and long-term evolutionary dynamics. Capitalizing on a large-scale screening of wild mammals, and genomic sequencing, we expand the novel rodent host range of hepaciviruses and document further virus diversity. We infer a significant influence of frequent cross-species transmission as well as some signal for virus-host co-divergence, and find comparative host and geographic structure. We also provide the first formal estimates of the timescale of hepaciviruses indicating an origin of about 22 million years ago. Our study offers new insights in hepacivirus evolutionary dynamics with broadly applicable methods that can support future research in virus evolution.

4
Host evolutionary history and ecology shape virome composition in fishes

Geoghegan, J. L.; Di Giallonardo, F.; Wille, M.; Ortiz-Baez, A. S.; Costa, V. A.; Ghaly, T.; Mifsud, J. C.; Turnbull, O. M.; Bellwood, D. R.; Williamson, J. E.; Holmes, E. C.

2020-05-08 microbiology 10.1101/2020.05.06.081505 medRxiv
Top 0.1%
75.8%
Show abstract

Revealing the determinants of virome composition is central to placing disease emergence in a broader evolutionary context. Fish are the most species-rich group of vertebrates and so provide an ideal model system to study the factors that shape virome compositions and their evolution. We characterised the viromes of 19 wild-caught species of marine fish using total RNA sequencing (meta-transcriptomics) combined with analyses of sequence and protein structural homology to identify divergent viruses that often evade characterisation. From this, we identified 25 new vertebrate-associated viruses and a further 22 viruses likely associated with fish diet or their microbiomes. The vertebrate-associated viruses identified here included the first fish virus in the Matonaviridae (single-strand, negative-sense RNA virus). Other viruses fell within the Astroviridae, Picornaviridae, Arenaviridae, Reoviridae, Hepadnaviridae, Paramyxoviridae, Rhabdoviridae, Hantaviridae, Filoviridae and Flaviviridae and were sometimes phylogenetically distinct from known fish viruses. We also show how key metrics of virome composition - viral richness, abundance and diversity - can be analysed along with host ecological and biological factors as a means to understand virus ecology. Accordingly, these data suggest that that the vertebrate-associated viromes of the fish sampled here are predominantly shaped by the phylogenetic history (i.e. taxonomic order) of their hosts, along with several biological factors including water temperature, habitat depth, community diversity and swimming behaviour. No such correlations were found for viruses associated with porifera, molluscs, arthropods, fungi and algae, that are unlikely to replicate in fish hosts. Overall, these data indicate that fish harbour particularly large and complex viromes and the vast majority of fish viromes are undescribed.

5
Evolution of ssDNA plant viruses in the natural environment - a journey through time

Herrera da Silva, J. P.; Xavier, C. A. D.; Oliveira, P. G. S.; Godinho, M. T.; Lage, J. B.; Silva, J. C. F.; Lima, A. T. M.; Zerbini, F. M.

2025-09-11 microbiology 10.1101/2025.09.10.675476 medRxiv
Top 0.1%
72.5%
Show abstract

Begomoviruses pose a major threat to food security, particularly in developing countries. These small ssDNA viruses exhibit substitution rates comparable to those of RNA viruses. The temporal dynamics of begomoviruses in non-agricultural environments have been largely overlooked, and little is known about how these viruses evolve in the absence of anthropogenic influence. In this study, we investigated the temporal dynamics of begomoviruses in a small fragment of regenerating Atlantic Forest area in the state of Minas Gerais, Brazil. Samples of Sida acuta, a wild plant native from South America, were collected at the same location over a 12-year period (2011-2022). Five distinct begomoviruses were detected infecting this host: OxYVV, SiYLCV, SimMV, MaYVV, and SiMV. OxYVV and SimMV were subdivided into multiple variants, revealing their potential as reservoirs of viral biodiversity. Several shifts in species and variant composition were observed in the viral community over time, with the most drastic change occurring in 2016, when SiYLCV outnumbered OxYVV. The reasons behind this turnover remain uncertain, but the most compelling clues point to a population expansion of SiYLCV. We detected a strong temporal signal in two of the most abundant viruses (OxYVV and SiYLCV), which allowed us to calibrate molecular clocks and estimate substitution rates for both of them. Our results indicate that, even in the natural environment, begomoviruses can evolve at rates similar to those reported in agricultural systems.

6
A comprehensive genealogy of the replication associated protein of CRESS DNA viruses reveals a single origin of intron-containing Rep

Zhao, L.; Lavington, E.; Duffy, S.

2019-07-01 evolutionary biology 10.1101/687855 medRxiv
Top 0.1%
72.4%
Show abstract

Abundant novel circular Rep-encoding ssDNA viruses (CRESS DNA viruses) have been discovered in the past decade, prompting a new appreciation for the ubiquity and genomic diversity of this group of viruses. Although highly divergent in the hosts they infect or are associated with, CRESS DNA viruses are united by the homologous replication-associated protein (Rep). An accurate genealogy of Rep can therefore provide insights into how these diverse families are related to each other. We used a dataset of eukaryote-associated CRESS DNA RefSeq genomes (n=926), which included representatives from all six established families and unclassified species. To assure an optimal Rep genealogy, we derived and tested a bespoke amino acid substitution model (named CRESS), which outperformed existing protein matrices in describing the evolution of Rep. The CRESS model-estimated Rep genealogy resolved the monophyly of Bacilladnaviridae and the reciprocal monophyly of Nanoviridae and the alpha-satellites when trees estimated with general matrices like LG did not. The most intriguing, previously unobserved result is a likely single origin of intron-containing Reps, which causes several geminivirus genera to group with Genomoviridae (bootstrap support 55%, aLRT SH-like support 0.997, 0.91-0.997 in trees estimated with established matrices). This grouping, which eliminates the monophyly of Geminiviridae, is supported by both domains of Rep, and appears to be related to our use of all RefSeq Reps instead of subsampling to get a smaller dataset. In addition to producing a trustworthy Rep genealogy, the derived CRESS matrix is proving useful for other analyses; it best fit alignments of capsid protein sequences from several CRESS DNA families and parvovirus NS1/Rep sequences.

7
Differences in codon usage between host-species-specific rabies virus clades are driven by UpA and purine content

Durrant, R.; Dushoff, J.; Arnold, M.; Cobbold, C.; Hampson, K.

2026-01-12 microbiology 10.64898/2026.01.12.699068 medRxiv
Top 0.1%
71.7%
Show abstract

Viral genes sometimes use certain codons more than others due to their nucleotide content, translational efficiency, and selection pressure from the host immune system. The rabies virus (RABV) is a negative strand RNA virus which can infect a broad range of mammalian hosts, with many of its clades circulating predominantly in specific host species. Previous work on codon usage in RABV has focused only on broader viral clades. We use publicly available RABV nucleoprotein gene sequences to investigate how dinucleotide content and codon usage biases differ between host-associated clades, and what drives these differences. We found that codon usage varies most between bat- and carnivore-associated RABV clades, and more subtly between host-species-specific minor clades within these groups. Pyrimidine and UpA content were both found to have a strong influence over codon usage patterns, and CpG content was considerably higher in carnivore-associated RABV clades than in bat-associated clades. This, along with a reduced number of zinc-finger antiviral protein binding motifs in bat-associated RABV sequences, suggests that bat-associated RABV clades may be under higher selection pressure from the hosts zinc-finger antiviral protein than carnivore-associated clades are, warranting further investigation of the mechanism underpinning this change.

8
Lineage-aware evolutionary analysis of hepatitis C virus within-host dynamics

Zhao, L.; Hall, M.; Giridhar, P.; Ghafari, M.; Kemp, S.; Chai, H.; Klenerman, P.; Barnes, E.; Ansari, M. A.; Lythgoe, K. A.

2024-10-17 evolutionary biology 10.1101/2024.10.15.617766 medRxiv
Top 0.1%
71.4%
Show abstract

Analysis of viral genetic data has previously revealed distinct within-host population structures in both untreated and interferon-treated chronic hepatitis C virus (HCV) infections. While multiple subpopulations persisted during the infection, each subpopulation was observed only intermittently. However, it was unknown whether similar patterns were also present after Direct Acting Antiviral (DAA) treatment, where viral populations were often assumed to go through narrow bottlenecks. Here we tested for the maintenance of population structure after DAA treatment failure. We analysed whole-genome next-generation sequencing data generated from a randomised study using DAAs (the BOSON study). We focused on samples collected from patients (N=84) who did not achieve sustained virological response (i.e. treatment failure) and had sequenced virus from multiple timepoints. For each individual, we tracked concordance in nucleotide variant frequencies through time. Using a sliding window approach, we applied sequenced-based and tree-based clustering algorithms across the entire HCV genome. Finally, we reconstructed viral haplotypes and estimated lineage specific within-host divergence rates from the haplotype phylogenies. Distinct viral subpopulations were maintained among a high proportion of individuals post DAA treatment failure. Using maximum likelihood modelling and model comparison, we found an overdispersion of viral evolutionary rates among individuals, and significant differences in evolutionary rates between lineages within individuals. These results suggest the virus is compartmentalised within individuals, with the varying evolutionary rates due to different viral replication rates or different selection pressures. We propose lineage awareness in future analyses of HCV evolution and infections to avoid conflating patterns from distinct lineages, and to recognise the likely existence of unsampled subpopulations.

9
Global diversity and dispersal routes of the Ostreid herpesvirus type 1 infecting Magallana gigas

Pelletier, C.; Chevignon, G.; Jacquot, M.; Morga, B.

2026-06-05 evolutionary biology 10.64898/2026.06.05.730385 medRxiv
Top 0.1%
70.6%
Show abstract

The order Herpesvirales comprises double-stranded DNA viruses characterized by substantial genomic plasticity, including recombination, structural variation, gene gain and loss, and lineage turnover. These processes can obscure phylogenetic relationships and complicate the reconstruction of viral evolutionary histories. Within this order, Ostreid herpesvirus 1 (OsHV-1) is a major pathogen of the Pacific oyster Magallana gigas and is responsible for recurrent mortality events affecting global aquaculture. Early molecular investigations based on partial genomic regions identified several viral lineages, including the "var" and "{micro}Var" lineages, but provided limited resolution for genome-wide evolutionary inference. The subsequent availability of complete genomes revealed extensive structural variation, such as insertions, deletions, and genomic rearrangements, highlighting the high genomic plasticity of OsHV-1. Although phylogenomic analyses have estimated evolutionary rates compatible with other large double-stranded DNA viruses, current inferences remain based on geographically restricted datasets, leaving the global evolutionary dynamics of OsHV-1 within its principal host insufficiently resolved. Here, we present 275 newly sequenced OsHV-1 genomes collected from infected M. gigas oysters between 1994 and 2022 across major oyster-producing regions worldwide. Using de novo genome assembly combined with comparative genomics, population genetic analyses, and time-scaled phylogenetic reconstruction, we investigate global genomic diversity and the spatio-temporal dynamics of viral diversification. Our results reveal long-standing viral diversity in East Asia, the emergence of structurally distinct Pacific and microvariants lineages, and ongoing diversification shaped by recombination, structural genome plasticity, and anthropogenic oyster movements. By integrating three decades of whole-genome data, this study provides a phylogenomic framework for understanding the diversity, evolution, and dispersal of OsHV-1 in modern aquaculture systems.

10
Diverse patterns of intra-host genetic diversity in chronically infected SARS-CoV-2 patients

Rutsinsky, N.; Ben Zvi, A.; Fabian, I.; Segev, S. T.; Jacobi, B.; Harari, S.; Meijer, S.; Paran, Y.; Stern, A.

2024-12-16 evolutionary biology 10.1101/2024.11.23.624482 medRxiv
Top 0.1%
69.7%
Show abstract

In rare individuals with a severely immunocompromised system, chronic infections of SARS-CoV-2 may develop, where the virus replicates in the body for months. Sequencing of some chronic infections has uncovered dramatic adaptive evolution and fixation of mutations reminiscent of lineage-defining mutations of variants of concern (VOCs). This has led to the prevailing hypothesis that VOCs emerged from chronic infections. To examine the mutation dynamics and intra-host genomic diversity of SARS-CoV-2 during chronic infections, we focused on a cohort of nine immunocompromised individuals with chronic infections and performed longitudinal sequencing of viral genomes. We show that sequencing errors may cause erroneous inference of high genetic diversity, and to overcome this we used duplicate sequencing across patients and time-points, allowing us to distinguish errors from low frequency mutations. We further find recurrent low frequency mutations that we flag as most likely sequencing errors. This stringent approach allowed us to reliably infer low frequency mutations and their dynamics across time. We inferred a synonymous divergence rate of the virus of [~]2x10-6 mutations/base/day, consistent with the SARS-CoV-2 mutation rate estimated in tissue culture. The rate of non-synonymous divergence varied widely among the different patients. We highlight two patients with opposing patterns: in one patient the rate of divergence was zero, yet this patient harbored multiple presumably defective viruses at low frequencies throughout the infection. Another patient exhibited dramatic adaptive evolution, including clonal competition. Overall, our results suggest that the emergence of highly divergent variants from chronic infections is likely a very rare event and this emphasizes the need to better understand the conditions that allow such emergence events.

11
Intrahost speciations and host switches shaped the evolution of herpesviruses

Brito, A. F.; Pinney, J. W.

2020-05-17 evolutionary biology 10.1101/418111 medRxiv
Top 0.1%
66.3%
Show abstract

Cospeciation has been suggested to be the main force driving the evolution of herpesviruses, with viral species co-diverging with their hosts along more than 400 million years of evolutionary history. Recent studies, however, have been challenging this assumption, showing that other co-phylogenetic events, such as intrahost speciations and host switches play a central role on their evolution. Most of these studies, however, were performed with undated phylogenies, which may underestimate or overestimate the frequency of certain events. In this study we performed co-phylogenetic analyses using time-calibrated trees of herpesviruses and their hosts. This approach allowed us to (i) infer co-phylogenetic events over time, and (ii) integrate crucial information about continental drift and host biogeography to better understand virus-host evolution. We observed that cospeciations were in fact relatively rare events, taking place mostly after the Late Cretaceous (~100 Millions of years ago). Host switches were particularly common among alphaherpesviruses, where at least 10 transfers were detected. Among beta- and gammaherpesviruses, transfers were less frequent, with intrahost speciations followed by losses playing more prominent roles, especially from the Early Jurassic to the Early Cretaceous, when those viral lineages underwent several intrahost speciations. Our study reinforces the understanding that cospeciations are uncommon events in herpesvirus evolution. More than topological incongruences, mismatches in divergence times were the main disagreements between host and viral phylogenies. In most cases, host switches could not explain such disparities, highlighting the important role of losses and intrahost speciations in the evolution of herpesviruses.

12
Machine Learning Reveals Key Glycoprotein Mutations and Rapidly Assigns Lassa Virus Lineages

Daodu, R. O.; Ulrich, J.-U.; Kuehnert, D.

2024-07-31 bioinformatics 10.1101/2024.07.31.605963 medRxiv
Top 0.1%
65.7%
Show abstract

Lassa fever, caused by the Lassa virus (LASV), remains a major public health concern in West Africa, causing numerous fatalities annually and several intercontinental cases since its discovery in 1969. Despite ongoing research, no approved vaccines are available, with current efforts focusing on immunotherapy. LASV is divided into distinct lineages that circulate in specific geographic regions, elicit varying immune responses, and exhibit different pathophysiological effects. Understanding the genetic differences between these lineages is crucial for developing, improving, and distributing diagnostics, treatments, and vaccines. In this study, we analyzed the LASV glycoprotein complex (GPC), the only surface protein, using statistics, machine learning, and phylogenetics. At a population scale, we identified key amino acid differences between Nigerian lineages and those in other West African regions, particularly near the stable signal peptide cleavage site and other immune-related regions (e.g., AA positions 59-76). Additionally, we found that GPC sequences from Lineages II and III are shorter than those from Lineage IV, due to the high prevalence of a codon insertion at positions 178-180 (amino acid position 60). This insertion may contribute to inaccuracies observed in molecular diagnostics for LASV and may also play a role in the increased fatality associated with Lineage IV. The insertion has reemerged and persisted in Lineage II which may indicate a fitness advantage. Furthermore, we developed a fast and highly accurate lineage classification tool called CLASV that allows rapid identification of LASV lineages, improving the ability to monitor emerging outbreaks and exported cases.

13
Co-evolutionary analysis suggests a role for TLR9 in papillomavirus restriction

King, K. M.; Larsen, B. B.; Gryseels, S.; Richet, C.; Kraberger, S.; Jackson, R.; Worobey, M.; Harrison, J. S.; Varsani, A.; Van Doorslaer, K.

2021-04-18 microbiology 10.1101/2021.04.17.440006 medRxiv
Top 0.1%
65.6%
Show abstract

A.Upon infection, DNA viruses can be sensed by pattern recognition receptors (PRRs) leading to the activation of type I and III interferons, aimed at blocking infection. Therefore, viruses must inhibit these signaling pathways, avoid being detected, or both. Papillomavirus virions are trafficked from early endosomes to the Golgi apparatus and wait for the onset of mitosis to complete nuclear entry. This unique subcellular trafficking strategy avoids detection by cytoplasmic PRRs, a property that may contribute to establishment of infection. However, as the capsid uncoats within acidic endosomal compartments, the viral DNA may be exposed to detection by toll-like receptor (TLR) 9. In this study we characterize two new papillomaviruses from bats and use molecular archeology to demonstrate that their genomes altered their nucleotide composition to avoid detection by TLR9, providing evidence that TLR9 acts as a PRR during papillomavirus infection. Furthermore, we demonstrate that TLR9, like other components of the innate immune system, is under evolutionary selection in bats, providing the first direct evidence for co-evolution between papillomaviruses and their hosts.

14
Genetic drift acts strongly on within-host influenza virus populations during acute infection but does not act alone

Shi, Y. T.; Martin, M. A.; Weissman, D. B.; Koelle, K.

2025-08-31 evolutionary biology 10.1101/2025.08.27.672713 medRxiv
Top 0.1%
65.3%
Show abstract

The evolutionary dynamics of seasonal influenza A viruses (IAVs) have been well characterized at the population level, with antigenic drift known to be a major force in driving strain turnover. The evolution of IAV populations at the within-host level, however, is still less well characterized. Improving our understanding of within-host IAV evolution has the potential to shed light on the source of new strains, including new antigenic variants, at the population level. Existing studies have pointed towards the role that stochastic processes play in shaping within-host viral evolution in acute infections of both humans and pigs. Here, we apply a population genetic model called the Beta-with-Spikes approximation to longitudinal intrahost Single Nucleotide Variant (iSNV) frequency data to quantify the extent of genetic drift acting on IAV populations at the within-host scale. We estimate small effective population sizes in both human IAV infections (NE = 41, 95% confidence interval: [22-72]) and swine IAV infections (NE = 10, 95% confidence interval: [8-14]). Moreover, we evaluate the consistency of the observed iSNV dynamics with Wright-Fisher model simulations. For the human IAV dataset that we analyze, we find that observed within-host IAV evolutionary dynamics are consistent with this classic model at the estimated low effective population size. However, for the swine IAV dataset, we find statistical evidence for rejecting the classic Wright-Fisher model as the only process governing within-host iSNV frequency dynamics. Our results contribute to the growing number of studies that point towards the important role of genetic drift in shaping patterns of genetic diversity in IAV populations within acutely infected hosts. It further raises questions about whether and what other processes, such as spatial compartmentalization, viral progeny production dynamics with strong skew, or selection, may be needed to explain patterns of within-host IAV evolution.

15
Transmission bottleneck size estimation from de novo viral genetic variation

Shi, T.; Harris, J.; Martin, M. A.; Koelle, K.

2023-08-14 evolutionary biology 10.1101/2023.08.14.553219 medRxiv
Top 0.1%
64.4%
Show abstract

Sequencing of viral infections has become increasingly common over the last decade. Deep sequencing data in particular have proven useful in characterizing the roles that genetic drift and natural selection play in shaping within-host viral populations. They have also been used to estimate transmission bottleneck sizes from identified donor-recipient pairs. These bottleneck sizes quantify the number of viral particles that establish genetic lineages in the recipient host and are important to estimate due to their impact on viral evolution. Current approaches for estimating bottleneck sizes exclusively consider the subset of viral sites that are observed as polymorphic in the donor individual. However, allele frequencies can change dramatically over the course of an individuals infection, such that sites that are polymorphic in the donor at the time of transmission may not be polymorphic in the donor at the time of sampling and allele frequencies at donor-polymorphic sites may change dramatically over the course of a recipients infection. Because of this, transmission bottleneck sizes estimated using allele frequencies observed at a donors polymorphic sites may be considerable underestimates of true bottleneck sizes. Here, we present a new statistical approach for instead estimating bottleneck sizes using patterns of viral genetic variation that arose de novo within a recipient individual. Specifically, our approach makes use of the number of clonal viral variants observed in a transmission pair, defined as the number of viral sites that are monomorphic in both the donor and the recipient but carry different alleles. We first test our approach on a simulated dataset and then apply it to both influenza A virus sequence data and SARS-CoV-2 sequence data from identified transmission pairs. Our results confirm the existence of extremely tight transmission bottlenecks for these two respiratory viruses, using an approach that does not tend to underestimate transmission bottleneck sizes.

16
Divergent hepaciviruses, chuvirus and deltaviruses in Australian marsupial carnivores (Dasyurids) identified through transcriptome mining

Harvey, E.; Mifsud, J. C.; Holmes, E. C.; Mahar, J. E.

2023-06-27 microbiology 10.1101/2023.06.27.546737 medRxiv
Top 0.1%
62.3%
Show abstract

Although Australian marsupials are characterised by unique biology and geographic isolation, little is known about the viruses present in these iconic wildlife species. The Dasyuromorphia are an order of marsupial carnivores found only in Australia that include both the extinct Tasmanian tiger (Thylacine) and the highly threatened Tasmanian devil. Several other members of the order are similarly under threat of extinction due to habitat loss, hunting, disease, and competition and predation by introduced species such as feral cats. We utilised publicly available RNA-seq data from the NCBI Sequence Read Archive (SRA) database to document the viral diversity within four Dasyuromorphia species. Accordingly, we identified 15 novel virus species from five DNA virus families (Adenoviridae, Anelloviridae, Herpesviridae, Papillomaviridae and Polyomaviridae) and three RNA virus taxa: the order Jingchuvirales, the genus Hepacivirus, and the delta-like virus group. Of particular note was the identification of a marsupial specific clade of delta-like viruses that may indicate an association of deltaviruses and with marsupial species dating back to their origin some 160 million years ago. In addition, we identified a highly divergent hepacivirus in a numbat liver transcriptome that falls outside of the larger mammalian clade, as well as the first detection of the Jingchuvirales in a mammalian host - a chu-like virus in Tasmanian devils - thereby expanding the host range beyond invertebrates and ectothermic vertebrates. As many of these Dasyuromorphia species are currently being used in translocation efforts to reseed populations across Australia, understanding their virome is of key importance to prevent the spread of viruses to naive populations.

17
Genomic analysis of Megalocytivirus genomes reveals widespread recombination

Hannaford, P. I.; Coff, L.; Hall, R.; Go, J.; Moody, N.; Lanfear, R. M.

2025-10-27 evolutionary biology 10.1101/2025.10.27.684589 medRxiv
Top 0.1%
60.7%
Show abstract

Megalocytiviruses are pathogens of global significance that can lead to substantial economic losses in aquaculture. Recombination among megalocytiviruses is typically assumed to be rare, although it has been relatively understudied. Here, we uncover widespread recombination within megalocytiviruses through detailed analyses of 63 Megalocytivirus genomes, including two which are newly sequenced and assembled. We also identify a number of genes which megalocytiviruses have likely obtained from outside the family Iridoviridae (iridovirids). These results have serious implications for the biosecurity management of megalocytiviruses, as they indicate that Megalocytivirus strains could be misclassified based on traditional approaches which target individual loci in the genome. We use this new knowledge of recombination to estimate updated phylogenetic trees of megalocytiviruses at the family-, genus-, and species-level. These trees show strong support for the designation of two novel species within the genus Megalocytivirus and highlight the difficulty of placing highly recombinant genomes in a single phylogenetic framework. We discuss the implications of our work for disease management, and the importance of genome-wide recombination detection and phylogenomic analysis in the classification and genetic characterisation of megalocytiviruses.

18
Insights into goatpox virus and sheeppox virus genomes from pangenome graphs

Downing, T.

2026-03-31 genomics 10.64898/2026.03.28.714820 medRxiv
Top 0.1%
60.3%
Show abstract

The Capripoxviruses (CaPV) comprise three species: goatpox virus (GTPV), sheeppox virus (SPPV) and lumpy skin disease virus (LSDV). They are large double-stranded DNA viruses with highly conserved core genomes and variable terminal regions. Previous studies have described variation in CaPV gene content, their broader population structure and the contribution of non-coding and structural variation remains opaque. This study investigated the genomic diversity and evolutionary history of GTPV and SPPV using an integrative framework combining phylogenetics, pangenome variation graphs (PVGs), and gene-specific analyses. We found marked differences in population structure between the two viruses. GTPV comprised three deeply divergent and genetically stable lineages with limited evidence of recent gene flow, whereas SPPV had weaker clade separation consistent with an ancestral bottleneck followed by recent population expansion. PVG-based analyses indicated that GTPV has a comparatively closed pangenome, while SPPV remains open, particularly at the genome termini. Structural and haplotypic variation was concentrated at the inverted terminal repeats (ITRs), which moderate host immunity and specificity. In several lineages, extended putative ORFs spanning adjacent terminal genes were observed, indicating recurrent structural plasticity at the genome ends. Patterns of gene-specific conservation and divergence highlighted loci under strong constraint and lineage-specific structural changes that may contribute to host specificity. Together, these results demonstrate how graph-based genome models complement gene-based analyses in resolving poxvirus genome evolution and provide a resource for improved comparative and population genomic studies of large DNA viruses. SignificanceCapripoxviruses are economically important livestock pathogens, yet the genomic mechanisms underlying their diversification and host specificity remain poorly resolved. By applying pangenome variation graphs alongside phylogenetic and gene-level analyses, this study reveals fundamental differences in how goatpox and sheeppox viruses have evolved. Goatpox virus had a deeper, more stable lineage structure, whereas sheeppox virus was more recent and diverse. Importantly, structural variation at the inverted terminal repeats emerged as a major driver of genomic diversity, including lineage-specific haplotypes and variable gene structures. These findings demonstrated the value of graph-based genome representations for resolving complex variation in large DNA viruses and provides a framework for improving genomic surveillance, comparative analyses, and future investigations into host range, virulence and tropism.

19
Characterizing a century of genetic diversity and contemporary antigenic diversity of N1 neuraminidase in IAV from North American swine

Hufnagel, D. E.; Young, K. M.; Arendsee, Z.; Gay, L. C.; Caceres, C. J.; Rajao, D. S.; Perez, D. R.; Vincent Baker, A. L.; Anderson, T. K.

2022-11-18 evolutionary biology 10.1101/2022.11.18.517097 medRxiv
Top 0.1%
60.1%
Show abstract

Influenza A viruses (IAV) of the H1N1 classical swine lineage became endemic in North American swine following the 1918 pandemic. Additional human-to-swine transmission events after 1918, and a spillover of H1 viruses from wild birds in Europe, potentiated a rapid increase in genomic diversity via reassortment between introductions and the endemic classical swine lineage. To determine mechanisms affecting reassortment and evolution, we conducted a phylogenetic analysis of N1 and paired HA swine IAV genes in North America between 1930 and 2020. We described fourteen N1 clades within the N1 Eurasian avian lineage (including the N1 pandemic clade) and the N1 classical swine lineage. Seven N1 genetic clades had evidence for contemporary circulation. To assess antigenic drift associated with N1 genetic diversity, we generated a panel of representative swine N1 antisera and quantified the antigenic distance between wild-type viruses using enzyme-linked lectin assays and antigenic cartography. Within the N1 lineage, antigenic similarity was variable and reflected shared evolutionary history. Sustained circulation and evolution of N1 genes in swine had resulted in significant antigenic distance between the N1 pandemic clade and classical swine lineage. We also observed a significant increase in the rate of evolution in the N1 pandemic clade relative to the classical lineage. Between 2010 and 2020, N1 clades and N1-HA pairings fluctuated in detection frequency across North America, with hotspots of diversity generally appearing and disappearing within two years. We also identified frequent N1-HA reassortment events (n = 36), which were rarely sustained (n = 6) and sometimes also concomitant with the emergence of new N1 genetic clades (n = 3). These data form a baseline from which we can identify N1 clades that expand in range or genetic diversity that may impact viral phenotypes or vaccine immunity and subsequently the health of North American swine.

20
Quantifying Asymmetric Coevolutionary Dynamics using Normalized Phylogenetic Costs

Wagle, S.; Markin, A.; Sherman, T. J.; Mayo, C.; Dunham, T. J.; Brelsfoard, C.; Cohnstaedt, L. W.; Wilson, W. C.; Anderson, T. K.; Eulenstein, O.

2026-07-03 bioinformatics 10.64898/2026.06.29.734822 medRxiv
Top 0.1%
59.7%
Show abstract

Coevolutionary studies aim to characterize associations, such as virus-host relationships, by using phylogenetic distances to quantify the topological concordance between the phylogenies of interacting taxa. However, phylogenetic distances cannot capture asymmetrical relationships that arise from differences in sampling, evolutionary rates, or characterizations between datasets. Furthermore, a lack of accurate normalization complicates the interpretation and validation of coevolutionary analyses. To address these limitations, we employed the Asymmetric Cluster Affinity and Cluster Support costs as a general framework to quantify coevolutionary patterns across multiple biological scales. We benchmarked the precision of these costs by reanalyzing a curated dataset documenting interspecies transmission frequencies across nineteen virus-host phylogenies. Our results corroborate prior findings showing that all virus families under study can cross species boundaries; however, the asymmetric costs provide a more granular representation, demonstrating that the frequency of such events varies significantly across families. We then applied the Asymmetric Cluster Support cost to quantify preferential gene segment pairings within the Bluetongue virus genome. This analysis revealed a close phylogenetic association between the outer capsid proteins VP2 and VP5, likely reflecting shared selective pressures due to their critical roles in cell entry and exit. In contrast, gene segments encoding nonstructural proteins exhibited discordant evolutionary histories relative to other segments. Finally, we demonstrated that the Asymmetric Cluster Support cost can detect coevolutionary dynamics in swine influenza A virus, identifying novel gene pairings indicative of major viral reassortment events. Overall, our approach demonstrates that normalized asymmetric phylogenetic costs accurately capture complex biological relationships and provide a robust framework for quantifying fine-scale coevolutionary dynamics in rapidly evolving pathogens.